Phase 2: Data & Mathematics Lesson 1 of 5

What is Data?
Structured vs Unstructured

Before you can build any AI system, you need to understand what it feeds on. Everything an AI knows comes from data. Let us understand exactly what data is, the different forms it takes, and how AI sees all of it as numbers.

You will learn
The difference between structured and unstructured data
How AI turns images, text and audio into numbers
What features and labels are in a dataset
How to read and describe any real dataset confidently

Data is the raw material of intelligence

Think about how a child learns to recognise a dog. Nobody writes them a manual that says "four legs, fur, barks." Instead, the child sees hundreds of dogs over several years: big ones, small ones, fluffy ones, spotted ones. Over time they build up an internal sense of what makes something a dog. Learning happens through exposure to examples.

AI works exactly the same way. You cannot just tell an AI model what a cat looks like. You have to show it thousands of photographs of cats. You cannot tell it what spam email sounds like. You have to give it tens of thousands of examples of spam and non-spam email. The examples are the data, and the data is everything.

This raises the obvious question: what exactly counts as data?

"In God we trust. All others must bring data."

W. Edwards Deming, statistician and engineer

The two main types of data

All data falls into one of two broad categories. Understanding this distinction is one of the first things you need to understand before choosing an AI approach for any problem.

Structured Data

Organised into rows and columns. Think of a spreadsheet or a database table. Each row is one record, each column is one attribute. It is tidy, labelled, and easy to search. Computers have been handling this kind of data for decades.

CSV files SQL databases Excel sheets Sensor readings Transaction records
Unstructured Data

Everything that does not fit neatly into rows and columns. The meaning is embedded inside the content itself, not in any predefined schema. It is harder to process but makes up the vast majority of data in the world.

Images & photos Audio recordings Video files Emails & text Social media posts

There is a third category worth knowing: semi-structured data. This sits in between. It is not a rigid table, but it does have some organisational tags or markers. JSON files and XML documents are good examples. A product listing in an online store might be semi-structured: it has defined fields like "price" and "category," but the product description is free-form text.

Stat worth knowing

Estimates suggest that around 80% of all data generated in the world is unstructured. This is why deep learning has become so important. It excels at finding patterns in raw images, audio and text in ways that classical machine learning methods simply cannot match.

A structured dataset up close

Let us look at a real example. The Titanic dataset, which you will use in your first coding exercise, is a classic structured dataset. Each row represents one passenger. Each column captures one fact about them.

PassengerId Survived Pclass Name Age Fare
103Braund, Mr. Owen227.25
211Cumings, Mrs. John3871.28
313Heikkinen, Miss. Laina267.93
411Futrelle, Mrs. Jacques3553.10
503Allen, Mr. William358.05

Notice a few important things. The Survived column is 0 or 1, not "yes" or "no." Computers and AI models do not naturally understand words. Everything eventually gets converted to numbers. The Name column is technically unstructured text sitting inside a structured table, and that kind of mixed reality is very common in the real world.

Features and labels: the vocabulary of ML

In machine learning, we have specific names for the different columns of a dataset. Understanding these terms will help you read papers, documentation and code.

📊
Features (inputs)

The columns that describe the thing you are analysing. In the Titanic dataset: age, ticket class, fare paid, number of siblings on board. These are what the model uses to make its prediction.

🎯
Labels (outputs)

The column you are trying to predict. In the Titanic dataset: Survived (0 or 1). In a house price dataset, it would be the actual price. This is the answer the model is learning to produce.

📦
Observations (rows)

Each individual record in your dataset. One row equals one observation. One passenger, one house, one transaction. The more observations you have, generally, the better your model can learn.

🔢
Dimensions

The number of features in your dataset. A dataset with 10 columns of features is 10-dimensional. High-dimensional data (hundreds or thousands of features) presents its own challenges, which we will encounter later.

How AI sees images: everything is numbers

Here is something that surprises most beginners. When you look at a photograph of a dog, you see a dog. When an AI model looks at the same photograph, it sees a grid of numbers.

A digital image is simply a grid of pixels. Each pixel has a colour, and that colour can be represented by three numbers: how much red, how much green and how much blue (the RGB values). Each value runs from 0 to 255. That means a 100 × 100 pixel image is actually a grid of 30,000 numbers (100 × 100 pixels × 3 colour channels).

8×8 grayscale image. Each cell is a pixel value (0 = black, 255 = white)

The model does not see a picture. It sees a matrix of numbers. Understanding this is the key to understanding how computers process images, faces, X-rays, satellite photos and more.

How AI sees text: tokens and numbers

Text cannot be fed directly into an AI model either. Before a model can process the sentence "The cat sat on the mat," it needs to be converted into numbers. This process is called tokenisation.

A tokeniser breaks text into chunks called tokens, usually words or parts of words. Each unique token gets assigned a number from a vocabulary list. So "cat" might become 4,823 and "sat" might become 7,112. The sentence becomes a list of numbers, which the model can then do mathematics on.

Think of it this way

Imagine a musician who can only read sheet music, not words. To communicate a story to them, you first have to translate every emotion and event into musical notes. The story is still there, just encoded in a language the musician can work with. AI models require exactly the same kind of translation, converting human-readable content into numbers the model can process.

Why data quality matters more than quantity

There is a common misconception that AI models simply need more data to get better. More is often better, but only if the data is good. The most important properties of a useful dataset are actually about quality, not size.

Relevant: Does the data actually contain signals that relate to the problem you are solving? Customer reviews are not useful data for predicting tomorrow's weather.

Representative: Does the dataset reflect the real world it is meant to represent? A model trained only on photos of faces with one skin tone will fail on faces with other skin tones. This is how bias enters AI systems.

Clean: Are there missing values, typos, duplicates, or impossible entries? Messy data produces unreliable models. In real AI projects, cleaning and preparing data consistently takes around 70–80% of the total time. You will experience this firsthand in Lesson 2.5.

Labelled (for supervised learning): For most machine learning tasks, someone has to label the data. This means a human tells the system which emails are spam, which photos contain cats, which loan applications defaulted. Labelling is expensive, time-consuming, and often underestimated.

Lesson Activity · No code required
Dataset Detective
This exercise builds the skill of looking at any dataset and immediately understanding its structure. This is something every AI practitioner does before writing a single line of code.
01 Go to kaggle.com/datasets (free account required) and find any dataset that interests you. Sports, music, health, movies. Any topic works.
02 Without downloading anything, answer these questions from the preview: How many rows? How many columns? Which columns are features and which might be a label? Are there any missing values visible?
03 Write a 3-sentence description of the dataset as if explaining it to a colleague. Bring it to the live session. The goal is developing the habit of asking the right questions before touching any data.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Pause & Reflect

Check your understanding

Click any question to reveal a thinking prompt. There are no wrong answers.

Your phone's camera stores a photo as a grid of pixel values. Is that structured or unstructured data? What about a voice recording?

A photo is a 3D matrix of numbers (height x width x colour channels) — technically structured in memory, but the semantic meaning (a face, a dog, a mountain) is unstructured. A voice recording is a sequence of amplitude values over time — structured as an array, unstructured in meaning. The key insight: computers store everything as numbers; "unstructured" describes meaning, not storage format.

A hospital wants to train an AI to predict patient readmission. What types of data would they need? Which would be structured, and which unstructured?

Structured: age, blood pressure readings, lab results (numbers in rows and columns), diagnosis codes, length of stay. Unstructured: doctor's written notes, discharge summaries, X-ray images, audio from consultations. The richest information is often in the unstructured parts — but it is the hardest for AI to process.

Why does AI generally perform better with more data? What are the limits of "more data is always better"?

More data exposes the model to more variation, helping it learn patterns that generalise. But the limits are real: more biased data makes the model more biased; more data costs more to store, label, and process; past data may not reflect the future; and after a certain point, returns diminish — a model trained on 100M examples often beats one trained on 1M, but rarely beats 1B by the same margin.

Progress
Done with this lesson?
Mark it complete to track your progress.